‹ BackNewsLLM review

LLM review

Google
2026-08-16 16:02:49

Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show

Google launched Gemini 3.7 Flash on August 13 and made it generally available in more than 160 countries on day one. According to Decrypt’s review, the model accepts up to 1 million input tokens, returns 64,000 output tokens, handles images, video, audio, and PDFs, and can use tools while operating a computer. Google’s own benchmark sheet says the model beats Claude Sonnet 5 and GPT-5.6 Terra in 11 of 18 tested categories, including 1,588 Elo on Code Arena’s web development board and 30.4% on AutomationBench, though Decrypt notes those figures come from Google’s methodology and should be treated as company claims rather than settled fact. Decrypt’s hands-on tests found the sharpest improvement in coding. Gemini 3.7 Flash generated a playable browser game on the first try in 2 minutes and 13 seconds, a major step up from Gemini 3.6 Flash, which Decrypt said could not produce a working file in a similar test after its July 21 release. Results were less convincing elsewhere. In creative writing, Decrypt said Gemini produced a tidy story but broke the central prompt rule, losing to a free community model, Qwopus3.5-27B-v3. In associative reasoning, logic, and advanced math, the review said Gemini often showed decent structure but failed on crucial task requirements, including a bridge puzzle and a polynomial problem it left unfinished. Decrypt’s conclusion: Gemini 3.7 Flash is a strong low-cost execution model inside Google’s ecosystem, but its creativity and reasoning remain uneven.

1560
Gemini 3.7 Flash review: big coding gains, weaker reasoning and writing still show
OpenAI
2026-07-18 15:21:03

GPT-5.6’s Sol, Terra, and Luna Put Fresh Pressure on Claude Fable 5

OpenAI’s GPT-5.6 rollout marks a structural shift: instead of one model with adjustable reasoning settings, the company released three distinct large language models—Sol, Terra, and Luna—with separate training profiles, pricing, and performance ceilings. In Decrypt’s review, the most meaningful matchup is Sol versus Claude Fable 5, Anthropic’s strongest public model. Sol is priced at $5 per million input tokens and $30 per million output tokens, while Fable 5 costs $10 and $50. The gap matters because Sol leads Fable 5 on several coding-focused benchmarks, and the cheaper Luna is already reported to beat Anthropic’s Opus 4.8 on coding. The pricing and product comparison gets sharper because Fable 5 has been operating under repeated access extensions. After a June 12 U.S. government ban tied to an Amazon researchers’ jailbreak finding, Anthropic withdrew the model globally for 19 days, restored it on July 1 with a new safety classifier, and then repeatedly delayed a planned shift to usage-credit billing. The latest deadline is July 19. Decrypt’s hands-on tests produced a split verdict. Fable 5 was judged better in creative writing and in a one-shot browser game build, while Sol performed better in readability-heavy tasks and led on multiple coding benchmarks. On broader intelligence scoring, the gap was minimal, with Fable 5 ahead by a single point on the cited aggregate index.

2170
GPT-5.6’s Sol, Terra, and Luna Put Fresh Pressure on Claude Fable 5